Cross-Tokenizer Challenges in RL

In RL, when the reference model and to-be-trained model have different tokenizers, training with RL algorithms will encounter some problems. Such problems are usually referred to as “cross-tokenizer” (cross-vocabulary) problems.

Date: 2026-09-30 Wed

Author: ArcaLunar